Top 100 Inference Engine Repositories
Ranking
| Ranking | Project Name | Stars | Forks | Language | Open Issues | Description | Last Commit |
|---|---|---|---|---|---|---|---|
| 1 | vllm | 90,927 | 21,676 | Python | 2325 | A high-throughput and memory-efficient inference and serving engine for LLMs | 2026-09-04 |
| 2 | ds4 | 22,054 | 2,065 | C | 240 | DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm | 2026-09-03 |
| 3 | web-llm | 18,956 | 1,373 | TypeScript | 134 | High-performance In-browser LLM Inference Engine | 2026-09-03 |
| 4 | ml-engineering | 18,890 | 1,236 | Python | 2 | Machine Learning Engineering Open Book | 2026-09-04 |
| 5 | MNN | 16,017 | 2,427 | C++ | 27 | MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI. | 2026-09-03 |
| 6 | Paddle-Lite | 7,273 | 1,619 | C++ | 46 | PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎) | 2026-04-27 |
| 7 | gemma.cpp | 7,035 | 659 | C++ | 26 | lightweight, standalone C++ inference engine for Google's Gemma models. | 2026-09-03 |
| 8 | cactus | 5,979 | 500 | C++ | 41 | Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots. | 2026-08-26 |
| 9 | shimmy | 5,825 | 561 | Rust | 9 | ⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary. | 2026-08-30 |
| 10 | DALI | 5,751 | 677 | C++ | 191 | A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications. | 2026-09-03 |
| 11 | CTranslate2 | 4,661 | 526 | C++ | 228 | Fast inference engine for Transformer models | 2026-08-31 |
| 12 | Tengine | 4,532 | 982 | C++ | 244 | Tengine is a lite, high performance, modular inference engine for embedded device | 2025-03-06 |
| 13 | Rapid-MLX | 3,649 | 413 | Python | 45 | The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replace... | 2026-09-04 |
| 14 | TransformerEngine | 3,519 | 818 | Python | 160 | A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance wi... | 2026-09-03 |
| 15 | spiceai | 3,075 | 226 | Rust | 717 | Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents. | 2026-09-04 |
| 16 | xDiT | 2,707 | 342 | Python | 52 | xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism | 2026-09-03 |
| 17 | h3.c | 2,555 | 193 | C | 26 | MiniMax H3 inference engine for Mac computers | 2026-08-11 |
| 18 | openlake | 2,405 | 422 | Rust | 111 | OpenLake is a high performance storage engine for efficient LLM inference and GPU Training | 2026-09-03 |
| 19 | AI-Engineering.academy | 2,378 | 277 | Jupyter Notebook | 6 | Mastering Applied AI, One Concept at a Time | 2026-02-27 |
| 20 | warp | 2,351 | 174 | C | 7 | Run the full 2.78-trillion-parameter Kimi K3 model or GLM-5.3-Flash beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine. | 2026-08-28 |
| 21 | gpu-perf-engineering-resources | 2,288 | 252 | Python | 0 | A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference. | 2026-08-23 |
| 22 | audio.cpp | 2,260 | 272 | C++ | 10 | An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependenc... | 2026-09-04 |
| 23 | tokenspeed | 2,088 | 274 | Python | 13 | TokenSpeed is a speed-of-light LLM inference engine. | 2026-09-04 |
| 24 | ai-performance-engineering | 1,907 | 263 | Python | 3 | Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning. | 2026-08-31 |
| 25 | sonar | 1,847 | 209 | C++ | 80 | Large-scale LLM inference engine | 2026-08-13 |
| 26 | Genie-TTS | 1,761 | 120 | Python | 31 | GPT-SoVITS ONNX Inference Engine & Model Converter | 2026-08-30 |
| 27 | uzu | 1,709 | 74 | Rust | 2 | A high-performance inference engine for AI models | 2026-09-03 |
| 28 | xllm | 1,555 | 290 | C++ | 83 | A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation. | 2026-09-04 |
| 29 | Atomic-Chat | 1,424 | 161 | TypeScript | 33 | Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V | 2026-09-03 |
| 30 | rtp-llm | 1,326 | 271 | Cuda | 38 | RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications. | 2026-09-04 |
| 31 | airunner | 1,314 | 103 | Python | 0 | Offline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflows | 2026-08-29 |
| 32 | Jlama | 1,303 | 164 | Java | 39 | Jlama is a modern LLM inference engine for Java | 2025-10-12 |
| 33 | cache-dit | 1,270 | 81 | Python | 85 | A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs. | 2026-09-02 |
| 34 | openrouter-runner | 1,258 | 123 | Python | 0 | Deprecated inference engine | 2025-09-06 |
| 35 | FeatherCNN | 1,227 | 275 | C++ | 18 | FeatherCNN is a high performance inference engine for convolutional neural networks. | 2019-09-24 |
| 36 | ezkl | 1,221 | 212 | Rust | 15 | ezkl is an engine for doing inference for deep learning models and other computational graphs in a zk-snark (ZKML). Use it from Python, Javascript, or the command line. | 2026-02-20 |
| 37 | tiny-vllm | 1,093 | 85 | C++ | 0 | Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM | 2026-08-23 |
| 38 | nobodywho | 1,091 | 77 | Rust | 11 | NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device. | 2026-09-03 |
| 39 | YOLOs-CPP | 1,081 | 162 | C++ | 0 | Cross-Platform Production-ready C++ inference engine for YOLO models (v5-v12, YOLO26). Unified API for detection, segmentation, pose estimation, OBB, and classification. Built on ONNX Runtime and Open... | 2026-08-23 |
| 40 | checkpoint-engine | 1,005 | 107 | Python | 3 | Checkpoint-engine is a simple middleware to update model weights in LLM inference engines | 2026-08-12 |
| 41 | ssd | 995 | 78 | Python | 2 | A lightweight inference engine supporting speculative speculative decoding (SSD). | 2026-05-10 |
| 42 | TinyChatEngine | 961 | 102 | C++ | 35 | TinyChatEngine: On-Device LLM Inference Library | 2024-07-04 |
| 43 | ZhiLight | 908 | 104 | C++ | 5 | A highly optimized LLM inference acceleration engine for Llama and its variants. | 2026-03-18 |
| 44 | kronk | 780 | 57 | Go | 7 | Your personal engine for running open source models locally. Use Go for hardware accelerated local inference with llama.cpp, whisper.cpp, and stablediffusion.cpp directly integrated into your Go appli... | 2026-09-03 |
| 45 | emlearn | 750 | 79 | Python | 16 | Machine Learning inference engine for Microcontrollers and Embedded devices | 2026-07-17 |
| 46 | atlas | 676 | 102 | Rust | 83 | Pure Rust Inference Engine | 2026-09-04 |
| 47 | pegainfer | 670 | 103 | Rust | 53 | Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2 | 2026-09-03 |
| 48 | libonnx | 652 | 114 | C | 16 | A lightweight, portable pure C99 onnx inference engine for embedded devices with hardware acceleration support. | 2026-07-07 |
| 49 | hipfire | 602 | 65 | Rust | 90 | RDNA-native LLM inference engine in Rust. | 2026-09-04 |
| 50 | tidy | 598 | 45 | Kotlin | 34 | Offline semantic Text-to-Image and Image-to-Image search on Android powered by quantized state-of-the-art vision-language pretrained CLIP model and ONNX Runtime inference engine | 2024-03-28 |
| 51 | swama | 592 | 32 | Swift | 38 | High-performance MLX-based LLM inference engine for macOS with native Swift implementation | 2026-09-04 |
| 52 | WhisperS2T | 578 | 75 | Jupyter Notebook | 31 | An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine | 2024-08-27 |
| 53 | qwen600 | 559 | 47 | Cuda | 0 | Static suckless single batch CUDA-only qwen3-0.6B mini inference engine | 2025-09-08 |
| 54 | FlashRT | 544 | 73 | C++ | 13 | FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Al... | 2026-08-31 |
| 55 | Anakin | 537 | 135 | C++ | 53 | High performance Cross-platform Inference-engine, you could run Anakin on x86-cpu,arm, nv-gpu, amd-gpu,bitmain and cambricon devices. | 2022-09-23 |
| 56 | VectorHub | 530 | 135 | Jupyter Notebook | 1 | Deprecated historical repo. Superlinked now develops SIE, a self-hosted inference engine for embeddings, reranking, OCR, extraction, and document processing. | 2026-08-31 |
| 57 | dotLLM | 514 | 60 | C# | 193 | LLM inference engine written in .NET | 2026-07-30 |
| 58 | OpenArc | 513 | 44 | Python | 11 | Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints. | 2026-09-03 |
| 59 | zinc | 511 | 19 | Zig | 4 | Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon | 2026-09-03 |
| 60 | simple-llm | 482 | 37 | Python | 0 | ~950 line, minimal, extensible LLM inference engine built from scratch. | 2026-01-09 |
| 61 | crabml | 470 | 45 | Rust | 24 | a fast cross platform AI inference engine 🤖 using Rust 🦀 and WebGPU 🎮 | 2025-01-04 |
| 62 | ntransformer | 465 | 19 | C++ | 2 | High-efficiency LLM inference engine in C++/CUDA. Run Llama 70B on RTX 3090. | 2026-02-22 |
| 63 | Crane | 462 | 53 | Rust | 26 | A Pure Rust based LLM, VLM, VLA, TTS, OCR Inference Engine, powering by Candle & Rust. Alternate to your llama.cpp but much more simpler and cleaner.. | 2026-08-31 |
| 64 | flash-tokenizer | 458 | 11 | C++ | 7 | EFFICIENT AND OPTIMIZED TOKENIZER ENGINE FOR LLM INFERENCE SERVING | 2026-02-02 |
| 65 | JetStream | 457 | 67 | Python | 14 | JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome). | 2026-01-05 |
| 66 | InfiniTensor | 447 | 74 | C++ | 24 | InfiniTensor is a high-performance inference engine tailored for GPUs and AI accelerators. Its design focuses on effective deployment and swift academic validation. | 2026-09-01 |
| 67 | gpu-rest-engine | 422 | 95 | C++ | 6 | A REST API for Caffe using Docker and Go | 2018-07-20 |
| 68 | TensorSharp | 406 | 39 | C# | 3 | A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It suppor... | 2026-09-04 |
| 69 | AutoGrad-Engine | 399 | 50 | C# | 0 | A complete GPT language model (training and inference) in ~600 lines of pure C#, zero dependencies | 2026-02-14 |
| 70 | StockInference-Spark | 382 | 194 | Java | 5 | Stock inference engine using Spring XD, Apache Geode / GemFire and Spark ML Lib. | 2016-06-03 |
| 71 | flex-nano-vllm | 358 | 21 | Python | 1 | FlexAttention based, minimal vllm-style inference engine for fast Gemma 2 inference. | 2025-11-02 |
| 72 | sentis-samples | 358 | 73 | C# | 11 | Inference Engine samples internal development repository. Contains example and template projects for Sentis package use. | 2026-08-12 |
| 73 | audio.cpp-webui | 348 | 73 | C++ | 3 | audio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generation. No Python dependency. | 2026-08-14 |
| 74 | rten | 333 | 26 | Rust | 42 | ONNX neural network inference engine | 2026-09-02 |
| 75 | AMDMIGraphX | 329 | 148 | C++ | 247 | AMD's graph optimization engine. | 2026-09-03 |
| 76 | RL-Kernel | 290 | 80 | Python | 81 | High-performance RL post-training infrastructure. Designed to achieve bitwise operator-level train-inference consistency across heterogeneous engines and extreme memory efficiency for GRPO, PPO, etc. | 2026-09-03 |
| 77 | elfi | 283 | 62 | Python | 10 | ELFI - Engine for Likelihood-Free Inference | 2025-05-07 |
| 78 | yolov4-triton-tensorrt | 282 | 61 | C++ | 3 | This repository deploys YOLOv4 as an optimized TensorRT engine to Triton Inference Server | 2022-06-02 |
| 79 | awesome-edge-machine-learning | 281 | 56 | Python | 1 | A curated list of awesome edge machine learning resources, including research papers, inference engines, challenges, books, meetups and others. | 2023-02-23 |
| 80 | dash-infer | 273 | 28 | C | 7 | DashInfer is a native LLM inference engine aiming to deliver industry-leading performance atop various hardware architectures, including CUDA, x86 and ARMv9. | 2025-08-06 |
| 81 | tflite2tensorflow | 272 | 42 | Python | 1 | Generate saved_model, tfjs, tf-trt, EdgeTPU, CoreML, quantized tflite, ONNX, OpenVINO, Myriad Inference Engine blob and .pb from .tflite. Support for building environments with Docker. It is possible ... | 2022-09-04 |
| 82 | whisper.el | 266 | 25 | Emacs Lisp | 8 | Speech-to-Text interface for Emacs using OpenAI's whisper model and whisper.cpp as inference engine. | 2026-07-17 |
| 83 | ai-hardware-engineer-roadmap | 265 | 39 | HTML | 0 | Master AI inference, AI agent harness systems, and hardware engineering — then design a physical AI chip. That is the goal. | 2026-09-03 |
| 84 | oramacore | 261 | 23 | Rust | 9 | OramaCore is the complete runtime you need for your projects, answer engines, copilots, and search. It includes a fully-fledged full-text search engine, vector database, LLM interface, and many more u... | 2026-04-14 |
| 85 | TurboLLM | 259 | 37 | TypeScript | 5 | Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first. | 2026-09-03 |
| 86 | compute-engine | 257 | 35 | C++ | 17 | Highly optimized inference engine for Binarized Neural Networks | 2026-07-30 |
| 87 | inferflow | 251 | 25 | C++ | 8 | Inferflow is an efficient and highly configurable inference engine for large language models (LLMs). | 2024-03-15 |
| 88 | lm-inference-engines | 241 | 9 | - | 8 | Comparison of Language Model Inference Engines | 2024-12-16 |
| 89 | KokoroSharp | 241 | 30 | C# | 11 | Fast local TTS inference engine in C# with ONNX runtime. Multi-speaker, multi-platform and multilingual. Integrate on your .NET projects using a plug-and-play NuGet package, complete with all voices. | 2026-08-14 |
| 90 | llm-inference-engineering | 240 | 28 | Markdown | 0 | Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs. | 2026-09-03 |
| 91 | Awesome-LLM-Inference-Engine | 237 | 22 | - | 1 | 2026-09-03 | |
| 92 | amd_inference | 234 | 8 | Python | 11 | Docker-based inference engine for AMD GPUs | 2024-10-07 |
| 93 | zse | 234 | 14 | Python | 1 | The inference engine the open-source world built for itself. | 2026-08-02 |
| 94 | MIVisionX | 217 | 92 | C++ | 12 | AMD MIVisionX is a computer vision toolkit built around a highly optimized, conformant open-source implementation of the Khronos OpenVX™ 1.3.2 specification. As of the 4.0.0 release, MIVisionX ships t... | 2026-09-03 |
| 95 | pulsar | 210 | 28 | Rust | 6 | SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places att... | 2026-09-01 |
| 96 | mlsub | 204 | 21 | OCaml | 11 | Prototype type inference engine | 2025-01-31 |
| 97 | embedded-ai.bench | 202 | 29 | Python | 17 | benchmark for embededded-ai deep learning inference engines, such as NCNN / TNN / MNN / TensorFlow Lite etc. | 2021-02-18 |
| 98 | rf-detr-cpp | 201 | 21 | C++ | 0 | Production-ready C++/TensorRT inference engine for RF-DETR. Object detection and instance segmentation with FP32/FP16/INT8 support. Optimized for NVIDIA GPUs, Jetson (Orin, AGX Thor). | 2026-08-14 |
| 99 | llm-systems-engineering-roadmap | 193 | 26 | - | 0 | A practical roadmap for mastering LLM internals, training, inference, RAG, agents, evaluation, and production architecture. | 2026-07-27 |
| 100 | microflow-rs | 190 | 30 | Rust | 3 | A robust and efficient TinyML inference engine. | 2026-05-26 |